Tutorials, deep dives and product notes — built for developers.
Terminal-Bench 4.0 checked Oct 6, 2026: Claude Opus 5.5 (64.8%) and Sonnet 5.5 (61.8%) now top the public tbench.ai board, with GPT-6 Astra and GPT-6.1 Sol tied at 58.2%. Vendor runs and Artificial Analysis results are reported separately, with harness, tokens and cost per row.
FrontierCode v1.1 Main: Opus 5.5 leads Cognition's board at 54.6%; Sonnet 5.5 adds 52.1% (xhigh, $1.59/rollout). Harness, per-run costs and API rates are distinguished. Updated October 2, 2026.
DeepSWE v1.1 leaderboard: Muse Spark 1.3 leads Datacurve at 75.4%. New separate provider runs: GPT-6.1 Sol 75.2%, Opus 5.5 74.2%, and Sonnet 5.5 71.0%; harnesses and unavailable per-task costs are disclosed.
Terminal-Bench 2.1: DeepSeek V4.1 Flash remains the public-board leader at 90.6%. Adds separate Vals AI snapshot runs: Opus 5.5 87.6% and Sonnet 5.5 83.1%, plus Cognition's unranked SWE-2 provider score.
SWE-bench Pro leaderboard: Anthropic-reported Opus 5.5 leads at 89.9%; Sonnet 5.5 scores 81.3%. Both results are provider runs, distinct from the standardized Scale leaderboard. Updated October 2, 2026.